Papers with large-scale standardized benchmarking of Norwegian generative language models
NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark (2025.findings-acl)
Copied to clipboard
Vladislav Mikhailov, Tita Enstad, David Samuel, Hans Christian Farsethås, Andrey Kutuzov, Erik Velldal, Lilja Øvrelid
| Challenge: | NorEval is a new evaluation suite for large-scale standardized benchmarking of Norwegian generative language models (LMs). |
| Approach: | They propose a new evaluation suite for large-scale standardized benchmarking of Norwegian generative language models (LMs) NorEval consists of 24 high-quality human-created datasets, of which five are created from scratch. |
| Outcome: | The evaluation framework and materials are publicly available. |